Papers by Tiago Timponi Torrent
Domain Adaptation in Neural Machine Translation using a Qualia-Enriched FrameNet (2022.lrec-1)
Copied to clipboard
| Challenge: | Neural models have been advancing in a myriad of tasks, but there is a lack of large training data. |
| Approach: | They propose a method for domain adaptation of Neural Machine Translation systems using a multilingual FrameNet enriched with qualia relations as an external knowledge base. |
| Outcome: | The proposed system outperforms the state-of-the-art commercial system in an experiment . the proposed system substitutes domain-specific terms in the source language by their adequate translation in the target language. |
Framed Multi30K: A Frame-Based Multimodal-Multilingual Dataset (2024.lrec-main)
Copied to clipboard
Marcelo Viridiano, Arthur Lorenzi, Tiago Timponi Torrent, Ely E. Matos, Adriana S. Pagano, Natália Sathler Sigiliano, Maucha Gamonal, Helen de Andrade Abreu, Lívia Vicente Dutra, Mairon Samagaio, Mariane Carvalho, Franciany Campos, Gabrielly Azalim, Bruna Mazzei, Mateus Fonseca de Oliveira, Ana Carolina Luz, Livia Padua Ruiz, Júlia Bellei, Amanda Pestana, Josiane Costa, Iasmin Rabelo, Anna Beatriz Silva, Raquel Roza, Mariana Souza Mota, Igor Oliveira, Márcio Henrique Pelegrino de Freitas
| Challenge: | Recent advances in image-captioning datasets combine image and language to solve a diverse range of tasks. |
| Approach: | They propose a Brazilian Portuguese multimodal-multilingual dataset that extends the Multi30K dataset with 158,915 original Brazilian Portuguese descriptions and 30,104 Brazilian Portuguese translations. |
| Outcome: | The proposed dataset adds 2,677,613 frame evocation labels to the 158,915 English descriptions and to the ones created for Brazilian Portuguese. |
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation (2025.findings-emnlp)
Copied to clipboard
Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Vladimir Araujo, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal, Jan Christian Blaise Cruz, Mary Dabre, Raj Dabre, Toqeer Ehsan, Naome A Etori, Fauzan Farooqui, Jiahui Geng, Guido Ivetta, Thanmay Jayakumar, Soyeong Jeong, Zheng Wei Lim, Aishik Mandal, Sofía Martinelli, Mihail Minkov Mihaylov, Daniil Orel, Aniket Pramanick, Sukannya Purkayastha, Israfel Salazar, Haiyue Song, Tiago Timponi Torrent, Debela Desalegn Yadeta, Injy Hamed, Atnafu Lambebo Tonja, Thamar Solorio
| Challenge: | a human-curated benchmark of over 5,800 triples of images is used to evaluate multimodal translation systems. |
| Approach: | They introduce a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. |
| Outcome: | The results show that visual context improves translation quality in culturally-specific items . |
Frame Shift Prediction (2022.lrec-1)
Copied to clipboard
| Challenge: | Frame shift is a cross-linguistic phenomenon in translation which results in corresponding pairs of linguistic material evoking different frames. |
| Approach: | They propose a task to predict cross-linguistic frame-to-frame correspondence and propose auxiliary training to learn cross-lingual frame-by-frame correlation. |
| Outcome: | The proposed task can learn cross-linguistic frame-to-frame correspondence and predict frame shifts in a Berkeley FrameNet-like configuration. |
Frame2: A FrameNet-based Multimodal Dataset for Tackling Text-image Interactions in Video (2024.lrec-main)
Copied to clipboard
Frederico Belcavello, Tiago Timponi Torrent, Ely E. Matos, Adriana S. Pagano, Maucha Gamonal, Natalia Sigiliano, Lívia Vicente Dutra, Helen de Andrade Abreu, Mairon Samagaio, Mariane Carvalho, Franciany Campos, Gabrielly Azalim, Bruna Mazzei, Mateus Fonseca de Oliveira, Ana Carolina Loçasso Luz, Lívia Pádua Ruiz, Júlia Bellei, Amanda Pestana, Josiane Costa, Iasmin Rabelo, Anna Beatriz Silva, Raquel Roza, Mariana Souza, Igor Oliveira
| Challenge: | et al., 2016) describe a multimodal dataset built from a Brazilian travel TV show . frameNet is composed of frames and their associated roles in a network of typed frame-to-frame relations. |
| Approach: | They present a multimodal dataset built from a Brazilian travel TV show annotated for FrameNet categories for both text and image communicative modes. |
| Outcome: | The proposed dataset includes 230 minutes of video annotated for FrameNet categories . the model can be applied to other communicative modes, i.e., images . |
SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models (2025.naacl-long)
Copied to clipboard
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Doughman, Ritam Dutt, Avijit Ghosh, Jessica Zosa Forde, Carolin Holtermann, Lucie-Aimée Kaffee, Tanmay Laud, Anne Lauscher, Roberto L Lopez-Davila, Maraim Masoud, Nikita Nangia, Anaelia Ovalle, Giada Pistilli, Dragomir Radev, Beatrice Savoldi, Vipul Raheja, Jeremy Qin, Esther Ploeger, Arjun Subramonian, Kaustubh Dhole, Kaiser Sun, Amirbek Djanibekov, Jonibek Mansurov, Kayo Yin, Emilio Villa Cueva, Sagnik Mukherjee, Jerry Huang, Xudong Shen, Jay Gala, Hamdan Al-Ali, null Tair Djanibekov, Nurdaulet Mukhituly, Shangrui Nie, Shanya Sharma, Karolina Stanczak, Eliza Szczechla, Tiago Timponi Torrent, Deepak Tunuguntla, Marcelo Viridiano, Oskar Van Der Wal, Adina Yakefu, Aurélie Névéol, Mike Zhang, Sydney Zink, Zeerak Talat
| Challenge: | Large Language Models reproduce and exacerbate social biases present in training data, and resources to quantify this issue are limited. |
| Approach: | They propose a multilingual parallel dataset to examine culturally-specific stereotypes that may be learned by LLMs. |
| Outcome: | The proposed dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. |